fix(kanban): judge goal deliverables before completion; fail open for operator errors - #960
Conversation
…dge errors Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected.
…surfaces Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class). - goals.kanban_handoff_rejection is now the single predicate (judge with completion_handoff=True; owned worker retries then blocks transient on judge error; operator fails open with a judge_error event; caller's conn). - Both surfaces' complete + request-review delegate to it, injecting only their run-id resolver and judge-availability probe. - CLI tests drive the real `kanban complete` / `request-review` argv path (build_parser -> kanban_command): completion_handoff reaches the judge and the card closes; real judge prompt accepts first completion; owned-worker 500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated. - AST contract: exactly one function in the tree calls the judge with completion_handoff, and both surface wrappers delegate to it. Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9 wiring/duplicate-predicate mutants).
|
🤖 merged-by: apollo · lane: discord · gate: BYPASS: FR paused by Ace ruling 2026-09-22; gate = kanban Argus PASS run 8602 · why: goal-mode judge: judge proposed deliverables, not a prior completion receipt (t_c4e23682): Argus PASS-WITH-CAVEATS r2 run 8602 at ddf8108 |
…sed replays reported lost (t_43e058b7) (#961) * fix: retain session prompt across route metadata writes (#958) Verified targeted session-state, restore, accounting, and model-resume tests: 236 passed, 1 pre-existing dashboard-auth fixture warning deselected. Reproduced NULL from billing route before fix. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(kanban): judge goal deliverables before completion; fail open for operator errors (#960) * fix(kanban): grade goal deliverables before completion and isolate judge errors Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected. * refactor(kanban): one shared goal-mode handoff gate for CLI and tool surfaces Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class). - goals.kanban_handoff_rejection is now the single predicate (judge with completion_handoff=True; owned worker retries then blocks transient on judge error; operator fails open with a judge_error event; caller's conn). - Both surfaces' complete + request-review delegate to it, injecting only their run-id resolver and judge-availability probe. - CLI tests drive the real `kanban complete` / `request-review` argv path (build_parser -> kanban_command): completion_handoff reaches the judge and the card closes; real judge prompt accepts first completion; owned-worker 500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated. - AST contract: exactly one function in the tree calls the judge with completion_handoff, and both surface wrappers delegate to it. Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9 wiring/duplicate-predicate mutants). --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(kanban): resume dependency-wait PR and page stranded ready cards (#952) * fix(kanban): respawn guard honors worker dependency_wait->promoted resume; add requeue_task + stuck-guard probe t_7d7ff489. Rule 4 active_pr no longer strands a card whose own dependency block (kind=dependency) postdates the newest PR comment and whose promotion has not yet spawned. requeue_task emits operator-intent 'requeued' for READY cards. respawn_guard_stuck_tasks lists cards held by active_pr >= 30 min. * fix(kanban): surface guarded ready cards and provide requeue verb Verify dependency_wait promotion dispatches once and subsequent crash is guarded; CLI requeue and one-shot alert tests pass (168 passed, 1 skipped in focused suites). t_7d7ff489. * test(kanban): keep corruption probe independent of watcher call count Verified: corrupt-board regression 2 passed, 22 deselected; ruff and diff check pass. Original two failures reproduced on clean fork base. * fix(kanban): keep PR continuation through status comments and unobserved ticks Verified 213 passed, 2 skipped across focused DB/CLI/watcher suites; subprocess stdin guard passed. * fix(kanban): consume event-ordered PR requeue intent Verified 183 passed, 1 skipped across DB/CLI/watcher; ruff and subprocess stdin guard pass. Same-second event-order mutant fails the regression test. * docs(kanban): state one-shot PR intent ordering * fix(kanban): bind PR comments to events and consume dependency intent Verified real dispatcher regressions RED before fix, then 210 passed, 2 skipped in focused DB/CLI/watcher suite; ruff and subprocess guard passed. * fix(kanban): pin READY requeue to PR comment identity Legacy same-second inline audit comments can share author and length; requeue snapshots the PR row id and remains one-shot. Verified 211 passed, 2 skipped; ruff clean. * fix(kanban): snapshot comment identity for every resume intent Verified focused DB/CLI/watcher/core suite: 217 passed, 2 skipped. Legacy equal-second strict mutant fails the intended arm. * fix(kanban): guard-stuck age ignores data events; board-scoped recovery command - respawn_guard_stuck_tasks: only kinds that can change the active_pr answer (_RESPAWN_GUARD_FAILURE_RESET_KINDS + dependency_wait + spawned, or a guard decline for another reason) restart the continuous-guard age; comments, heartbeats, attachments are data. - render_operator_command(board, verb, *args): single renderer, always emits --board <slug>; clear_verb uses it; watcher passes the probed board. - test_triage_resolve_records_who_and_why: expect after_comment_id == max(task_comments.id) at emit time (CI slice 15/16 red). Verified: new tests RED on 581299d (4 failed), GREEN here; recovery command executed via real CLI on default and secondary boards (rc0); no-board renderer mutant fails; 244 passed/1 skipped on db/triage/ watchers/cli/boards; ruff clean. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) (#966) * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) Card t_d3963974. Third Apollo boot-cost freeze. The cause: _fts_needs_rebuild_structural ran `SELECT COUNT(*) FROM messages` (SCAN messages USING COVERING INDEX, 2.5M rows) plus `COUNT(*) FROM messages_fts_docsize` on EVERY MessageStore/SummaryDAG construction, under plugins.context_engine._LOAD_LOCK. Measured on an APFS clone of the fleet DB: 13.4 s of a 17.9 s cold MessageStore() init. The two autocommit COUNTs could also straddle a concurrent ingest. They gave a false mismatch 15/300 times, and each mismatch triggered a full inline FTS drop+rebuild. That happened twice on 2026-09-24: held 779.8 s, 12 turns queued. Provenance: the count came in with the original vendor import 8b86963 (2026-06-16) and was carried unchanged through re-vendor 27b6178. #887/#902/#903 did not touch it. #903's plan check exempted USING COVERING INDEX and skipped messages_fts*, so it passed on this code. Fix: - The parity check moves to _fts_count_parity_mismatch. It never runs on the throttle=True (load) path. A metadata marker (fts_parity_checked_at:<fts>) throttles it to once per LCM_FTS_PARITY_CHECK_INTERVAL_HOURS (default 6 h). When due, it runs on the existing background integrity thread (own connection). A mismatch sets the /lcm doctor integrity flag instead of rebuilding inline. Both counts are read in one snapshot. - Explicit repair (throttle=False) and /lcm doctor still run parity synchronously. - A no-op load writes nothing: - _clear_integrity_failed runs only after a real repair. Before, it also erased the background corruption flag on every load. - The messages_dedup_v1 and schema_version upserts are marker-gated. - The integrity claim uses a 1 s busy timeout, not 30 s. - Lifecycle GC (on_session_start, every agent init): - no longer holds BEGIN IMMEDIATE across two SELECT DISTINCT session_id full scans; - uses indexed per-session probes; - runs at most once per 6 h per process. - _backfill_search_content no longer rewrites NULL over NULL for undecryptable rows. That rewrite fired msg_fts_update on every boot. Test: test_lcm_init_cost_regression now traces engine construction + on_session_start on the loading thread. It fails on ANY SCAN of messages, messages_fts*, summary_nodes and nodes_fts*, covering index included (only LIMIT-bounded statements are exempt). It also asserts that a steady-state load needs no write lock and preserves the corruption flag. RED on df43599 (3 failed); GREEN with this change; tests/context_engine: 393 passed. * test(lcm): lock parity-race repro and document restart meltdown --------- Co-authored-by: Apollo <apollo@angventures.io> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(gateway): reject torn spool keys and gate replay refusals Verified: 57 focused restart tests passed; forged invalid-key, tampered-admission, and two refusal-site arms exercised via real restart. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> Co-authored-by: Apollo <apollo@angventures.io>
FleetReviewReview: post-merge · head Post-merge review ( profile: light (rule: default light: lines 453<800, files 5<1000000, hunks 14<1000000, no hot path) · round 0 · members: B-state, L6, C-assert-xhigh, G · families: anthropic,openai,xai Confidence: 3/5 Findings
FleetReview provenance · models: B=gpt-6-sol, C=claude-code-opus-5-5, D=grok-4.6 · cost: $2.40 · duration: 11m 14s · rounds: 1 · files examined: 5 |
Goal-mode completion now grades the proposed deliverables without requiring a prior completion receipt. Worker judge errors retry once then block transient; operator completions fail open with a judge_error event. Both tool and CLI handoffs covered. Verification: 81 targeted tests passed; one unrelated ACP-dependent test excluded (acp not installed). Mutant disabling the completion rubric makes the first-completion regression fail.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.